iT邦幫忙

2026 iThome 鐵人賽

DAY 17
0
Build on Google AI

給藥袋裝一張嘴:30 天用 Android 與 Google VLM 實作高齡語音用藥助手系列 第 17 篇

Day 17|打造堅固的防禦牆!用 Pydantic 與 Mypy 實作 API 輸入校驗與強型別防衛

  • 分享至 

  • xImage
  •  

✏️【本日實作紀錄:靜態型別檢查 Mypy、Pydantic 與輸入資料防禦校驗】

進入專案第 17 天,系統逐漸從「原型驗證」邁向「生產級穩定度」。在多模態 AI 系統中,使用者上傳的藥袋圖片可能存在模糊、破損、解析度過低,甚至是非藥袋的無效照片;同時 API 接收到的 JSON 資料若格式不合規,容易導致後端處理邏輯拋出未捕獲的例外而崩潰。

今天我們將導入 Pydantic v2 資料模型驗證 與 Mypy 靜態型別檢查,為 PrescriptionVLM 建立嚴密的輸入資料防禦牆(Defensive Validation Guardrail),確保 API 輸入與 AI 模型輸出的資料結構皆具備百分之百的強型別保障。


一、防禦性程式設計(Defensive Programming)的核心價值

在結合 Vision-Language Models (VLM) 的 Web 服務中,資料防禦校驗有三大核心重點:

  • 輸入邊界防禦(Input Boundary Validation):在上傳圖片進入昂貴的 LLM API 呼叫前,先驗證檔案大小、MIME 類型與影像解析度,避免無效請求浪費 Token 與運算資源。
  • 結構化輸出校驗(Structured Output Validation):使用 Pydantic Model 替代原始 Dict 操作,強制確保 AI 傳回的 JSON 符合欄位型別要求,防止 KeyError 或 NoneType 錯誤。
  • 靜態型別安全(Static Type Safety):透過 Mypy 於編譯/部署前進行全專案程式碼型別檢查,及早發現潛在的型別不匹配 Bug。

二、定義 Pydantic 資料驗證模型(schemas.py)

我們建立獨立的 schemas.py 模組,專門管理所有 API 的 Request/Response 結構與 Pydantic 資料模型。

1. 安裝 Pydantic 與 Mypy 套件

確保 requirements.txt 包含以下依賴套件:

pydantic>=2.0.0
mypy>=1.5.0

2. 建立 schemas.py 檔案

在專案根目錄下新建 schemas.py:

from typing import List, Optional
from pydantic import BaseModel, Field, field_validator

class MedicineItem(BaseModel):
    """單一藥品明細驗證模型"""
    name: str = Field(..., description="藥品名稱", min_length=1)
    type: str = Field(..., description="藥品種類 (口服/外用)")
    frequency: str = Field(..., description="服用頻率 ex. 一天三次")
    dosage: str = Field(..., description="每次劑量 ex. 一顆")
    timing: str = Field(..., description="吃藥時機 ex. 飯後")
    warning: Optional[str] = Field(default="", description="單一藥品特別警語")

    @field_validator('type')
    @classmethod
    def validate_type(cls, v: str) -> str:
        allowed_types = ['口服', '外用', '其他']
        if not any(t in v for t in allowed_types):
            return '口服'  # 預設防呆回傳
        return v

class PrescriptionAnalysisResult(BaseModel):
    """藥單解析完整結果模型"""
    spoken_summary: str = Field(..., description="長輩白話播報摘要")
    safety_warnings: List[str] = Field(default_factory=list, description="用藥安全警語列表")
    medicines: List[MedicineItem] = Field(..., description="藥品清單")

class ImageValidationResponse(BaseModel):
    """圖片上傳校驗失敗回應"""
    error: str
    message: str
    detail: Optional[str] = None


三、整合 Pydantic 校驗與防衛邏輯至 app.py

將原本寫死的 JSON Schema 升級為 Pydantic 校驗,並在圖片讀取階段加入防衛檢查。

1. 開啟 app.py 並匯入 Pydantic 結構

from pydantic import ValidationError
from schemas import PrescriptionAnalysisResult, MedicineItem

2. 升級 analyze_prescription 路由防護

修改 app.py 中的圖片驗證與解析邏輯:

@app.route('/analyze-prescription', methods=['POST'])
@limiter.limit("5 per minute")
def analyze_prescription():
    start_time = time.time()
    
    # 1. 檔案存在性防護
    if 'image' not in request.files:
        logger.warning("請求缺少圖片檔案", extra={"extra_data": {"status_code": 400}})
        return jsonify({"error": "unsupported_media_type", "message": "未提供圖片檔案"}), 400

    file = request.files['image']
    
    # 2. 副檔名與檔名防護
    if file.filename == '' or not allowed_file(file.filename):
        logger.warning("上傳不支援的檔案格式", extra={"extra_data": {"filename": file.filename, "status_code": 400}})
        return jsonify({
            "error": "unsupported_media_type",
            "message": "不支援的檔案格式,請上傳 .jpg, .jpeg 或 .png 圖片。"
        }), 400

    try:
        # 3. 圖片邊界尺寸與損毀檢驗
        try:
            image = Image.open(file.stream)
            image.verify()  # 驗證圖片檔案完整性
            file.stream.seek(0)
            image = Image.open(file.stream)  # 重新開啟供後續處理
        except Exception as img_err:
            logger.warning(f"上傳圖片損毀或無效: {str(img_err)}")
            return jsonify({
                "error": "invalid_image",
                "message": "上傳的圖片檔案損毀或無法正常解碼,請重新拍照後再試。"
            }), 400

        prompt = """
        你是一位專業且細心的藥師助手。請分析這張藥袋照片:
        1. 將藥品分類為口服或外用,精準提取名稱、頻率、劑量與吃藥時間。
        2. 檢查是否有重複藥性或高風險注意事項,填入 safety_warnings。
        3. 針對高齡長輩,撰寫一段溫柔白話的 spoken_summary。
        """

        config = types.GenerateContentConfig(
            response_mime_type="application/json",
            response_schema=prescription_schema
        )
        
        vlm_start = time.time()
        model_name = os.getenv("GEMINI_MODEL", "gemini-3.8-flash")

        # 4. API 自動重試機制
        max_retries = 3
        response = None
        for attempt in range(max_retries):
            try:
                response = client.models.generate_content(
                    model=model_name,
                    contents=[image, prompt],
                    config=config
                )
                break
            except Exception as e:
                if ("503" in str(e) or "UNAVAILABLE" in str(e)) and attempt < max_retries - 1:
                    logger.warning(f"Gemini API 暫時忙碌 (503),等待 2 秒後第 {attempt + 2} 次重試...")
                    time.sleep(2)
                    continue
                raise e

        vlm_duration = round((time.time() - vlm_start) * 1000, 2)
        raw_json = json.loads(response.text)

        # 5. Pydantic 強型別結構校驗
        try:
            validated_data = PrescriptionAnalysisResult(**raw_json)
            result_data = validated_data.model_dump()
        except ValidationError as val_err:
            logger.error(f"Pydantic 模型校驗失敗: {val_err.json()}")
            return jsonify({
                "error": "validation_error",
                "message": "AI 解析成果格式不符規範,請重新嘗試上傳。",
                "detail": val_err.errors()
            }), 422

        spoken_text = result_data.get("spoken_summary", "解析完成。")
        safety_warnings = result_data.get("safety_warnings", [])

        if safety_warnings:
            prefix = "長輩請注意,這份藥單有特別需要留意的地方:" + ";".join(safety_warnings) + "。"
            spoken_text = f"{prefix} {spoken_text}"
            result_data['spoken_summary'] = spoken_text

        prescription_id = uuid.uuid4().hex[:8]
        filename = f"speech_{prescription_id}.mp3"
        filepath = os.path.join(AUDIO_DIR, filename)

        tts = gTTS(text=spoken_text, lang='zh-tw')
        tts.save(filepath)

        audio_url = f"/static/audio/{filename}"
        result_data['audio_url'] = audio_url
        result_data['prescription_id'] = prescription_id

        conn = sqlite3.connect(DATABASE_PATH)
        cursor = conn.cursor()

        warnings_json = json.dumps(safety_warnings, ensure_ascii=False)
        cursor.execute(
            "INSERT INTO prescriptions (id, spoken_summary, audio_url, safety_warnings) VALUES (?, ?, ?, ?)",
            (prescription_id, spoken_text, audio_url, warnings_json)
        )

        meds = result_data.get("medicines", [])
        for med in meds:
            cursor.execute(
                """INSERT INTO medicines 
                   (prescription_id, name, type, frequency, dosage, timing, warning) 
                   VALUES (?, ?, ?, ?, ?, ?, ?)""",
                (
                    prescription_id,
                    med.get("name"),
                    med.get("type"),
                    med.get("frequency"),
                    med.get("dosage"),
                    med.get("timing"),
                    med.get("warning", "")
                )
            )

        conn.commit()
        conn.close()

        send_line_notification(spoken_text, len(meds), safety_warnings, audio_url)

        total_duration = round((time.time() - start_time) * 1000, 2)
        logger.info(
            "藥單解析流程完成",
            extra={
                "extra_data": {
                    "prescription_id": prescription_id,
                    "vlm_duration_ms": vlm_duration,
                    "total_duration_ms": total_duration,
                    "medicines_count": len(meds),
                    "warnings_count": len(safety_warnings),
                    "status_code": 200
                }
            }
        )

        return jsonify(result_data), 200

    except Exception as e:
        total_duration = round((time.time() - start_time) * 1000, 2)
        logger.error(
            f"伺服器處理失敗: {str(e)}",
            extra={"extra_data": {"total_duration_ms": total_duration, "status_code": 500}}
        )
        return jsonify({"error": f"伺服器處理失敗: {str(e)}"}), 500


四、執行 Mypy 靜態型別檢查

先補下載套件:

pip install mypy pydantic

我們使用 Mypy 掃描專案,確認型別標註(Type Hints)完全符合規範。

1. 在 Terminal 執行 Mypy 檢查

mypy schemas.py --ignore-missing-imports

預期輸出:

Success: no issues found in 1 source file

五、測試與驗證

1. 重新編譯並執行 Docker 容器

docker rm -f prescription_service
docker build -t prescription-vlm:v1.0 .
docker run -d -p 5000:5000 --env-file .env --name prescription_service prescription-vlm:v1.0

2. 測試正常圖片上傳

curl -X POST http://127.0.0.1:5000/analyze-prescription \
  -F "image=@test_rx.jpg"

3. 測試非圖片損毀檔案上傳(防衛校驗驗證)

建立文字檔偽裝成圖片並上傳:

echo "not an image" > dummy.jpg
curl -X POST http://127.0.0.1:5000/analyze-prescription \
  -F "image=@dummy.jpg"

預期回傳:

{
  "error": "invalid_image",
  "message": "上傳的圖片檔案損毀或無法正常解碼,請重新拍照後再試。"
}

六、版本控制與提交 GitHub

測試完成後,將 Day 17 的修改提交至 GitHub:

git add .
git commit -m "保留雙引號 改填寫自己要記錄的標記 ex.鐵人賽第十七天"
git push

七、本日小結與明日預告

今天我們透過 Pydantic v2 與 Mypy,為專案構建了堅固的資料驗證防禦牆。無論是前端傳入的無效圖片檔,還是 LLM 產生的 JSON 欄位異常,系統都能第一時間精準捕獲並優雅提示,大幅提升了生產環境下的容錯度。

明日(Day 18),我們將進行 結構化日誌(Structured Logging)與觀測性系統升級,導入追蹤碼(Trace ID)與端到端效能監控,讓後端維運與除錯更加入微!


上一篇
Day 16|長輩看得到也聽得懂!高齡無障礙 UI/UX 與語音互動優化
系列文
給藥袋裝一張嘴:30 天用 Android 與 Google VLM 實作高齡語音用藥助手 共 17 篇
圖片
  熱門推薦
圖片
{{ item.channelVendor }} | {{ item.webinarstarted }} |
{{ formatDate(item.duration) }}
直播中

尚未有邦友留言

立即登入留言